Papers with multimodal baselines

8 papers
SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving (2026.findings-eacl)

Copied to clipboard

Challenge: Current models struggle to accurately decompose intricate visual inputs and connect perception with structured reasoning, leading to suboptimal performance.
Approach: They propose a Spatial Comprehension-Infused Symbolic Reasoning Framework to integrate spatial representations into structured symbolic reasoning chains.
Outcome: The proposed framework outperforms existing models in vision-intensive mathematical problems.
MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations (P19-1)

Copied to clipboard

Challenge: Emotion recognition in conversations has gained popularity due to its potential applications. Until now, a large multimodal multi-party emotional conversational database containing more than two speakers per dialogue was missing.
Approach: They propose to extend and enhance EmotionLines by combining 13,000 utterances from Friends dialogues with emotion and sentiment labels.
Outcome: The proposed dataset contains about 13,000 utterances from 1,433 dialogues from the TV-series Friends.
Open Your Model’s Eyes: Video and Context-Aware Multimodal Backchannel Prediction (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for predicting backchannels rely on audio and text . existing methods omit visual cues and conversational contexts for accurate prediction .
Approach: They propose a framework that leverages visual cues and conversational contexts to enhance backchannel prediction.
Outcome: The proposed framework outperforms existing methods and simple multimodal baselines in recognizing complex backchannels such as empathy.
MATCHED: Multimodal Authorship-Attribution To Combat Human Trafficking in Escort-Advertisement Data (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for human trafficking detection ignore the multimodal nature of online ads . sex trafficking is a pervasive crime exploiting individuals of all ages and genders .
Approach: They propose to use multimodal authorship attributes to identify suspicious ads that combine text and images to improve vendor identification and verification tasks.
Outcome: The proposed model outperforms existing methods for vendor identification and verification tasks using text-only, vision-only and multimodal training objectives.
CI-AVSR: A Cantonese Audio-Visual Speech Datasetfor In-car Command Recognition (2022.lrec-1)

Copied to clipboard

Challenge: In-car smart assistants should be able to process general as well as car-related commands and perform corresponding actions, which eases driving and improves safety.
Approach: They propose a dataset for in-car command recognition in the cantonese language with both video and audio data.
Outcome: The proposed model can achieve a considerable quality on the clean test set, but the speech recognition quality on noisy data is still inferior.
Not all Fake News is Written: A Dataset and Analysis of Misleading Video Headlines (2023.emnlp-main)

Copied to clipboard

Challenge: Social media platforms are used by half of U.S. adults for everyday news consumption.
Approach: They propose to analyze video headlines and whether annotators believe the headline is representative of the video’s contents.
Outcome: The proposed dataset analyzes video headlines and explains why annotators view a video as misleading.
MEXA: Towards General Multimodal Reasoning with Dynamic Multi-Expert Aggregation (2025.findings-emnlp)

Copied to clipboard

Challenge: MEXA is a training-free framework that performs modality- and task-aware aggregation of multiple expert models to enable effective multimodal reasoning across diverse domains.
Approach: MEXA is a training-free framework that performs modality- and task-aware aggregation of multiple expert models.
Outcome: MEXA performs modality- and task-aware aggregation of multiple expert models . it generates interpretable textual reasoning outputs and reasons over them using a Large Reasoning Model (LRM) MEX A consistently delivers performance improvements over strong multimodal benchmarks .
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing MLLMs are strong at understanding single plots, but struggle with multi-step reasoning . Existing approaches to manage context in chart reasoning include text-based chain-of-thought prompting .
Approach: They propose a hierarchical visual agent framework that iteratively constructs a working context in an image–text space.
Outcome: The proposed framework improves on strong multimodal baselines.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations